Statistical analysis of filled pauses’ rhythm for disfluent speech synthesis
نویسندگان
چکیده
Given that state of the art speech synthesis systems have already reached a high naturalness level, it is time to move to talking speech from the actual read speech framework. For this purpose it is thus necessary to investigate how disfluencies can be included in speech synthesis and even increase its naturalness. This paper builds on a previously presented work and focuses on finding a local model of filled pauses rhythm. A statistical study of rhythm effects around filled pauses is presented and based on the correlation between rhythm variables, a regression model is proposed to predict filled pauses duration and prepausal lengthening.
منابع مشابه
Statistical analysis of filled pauses2 rhythm for disfluent speech synthesis
Given that state of the art speech synthesis systems have already reached a high naturalness level, it is time to move to talking speech from the actual read speech framework. For this purpose it is thus necessary to investigate how disfluencies can be included in speech synthesis and even increase its naturalness. This paper builds on a previously presented work and focuses on finding a local ...
متن کاملDisfluent Speech Analysis and Synthesis: a preliminary approach
Despite of the existence of high quality unit selection speech synthesizers, they are based on a reading style approach. However, new applications such as Speech-to-Speech Translation or Speech User Interfaces demand a talking style which is more natural in these contexts. Disfluencies are a major characteristic of talking style so that it is convenient to be able to generate disfluent speech. ...
متن کاملBreath and Non-breath Pauses in Fluent and Disfluent Phases of German and French L1 and L2 Read Speech
In this study we examined the read speech of native and nonnative speakers with respect to pausing details of audible breathing, particularly in disfluent phases. 20 German and 20 French native speakers read the same narrative text in their native (L1) and in their non-native language (L2). Some expected results were confirmed: more frequent pauses and more frequent disfluencies in L2, as well ...
متن کاملA Lattice-based Approach to Automatic Filled Pause Insertion
This paper describes a novel method for automatically inserting filled pauses (e.g., UM) into fluent texts. Although filled pauses are known to serve a wide range of psychological and structural functions in conversational speech, they have not traditionally been modelled overtly by state-of-the-art speech synthesis systems. However, several recent systems have started to model disfluencies spe...
متن کاملSynthesising Filled Pauses: Representation and Datamixing
Filled pauses occur frequently in spontaneous human speech, yet modern text-to-speech synthesis systems rarely model these disfluencies overtly, and consequently they do not output convincing synthetic filled pauses. This paper presents a text-to-speech system that is specifically designed to model these particular disfluencies more efffectively. A preparatory investigation shows that a synthet...
متن کامل